European Journal of Epidemiology
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match European Journal of Epidemiology's content profile, based on 43 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Liu, W.; Collister, J.; Clifton, L.; Littlejohns, T.; Hunter, D. J.
Show abstract
1.Disease-specific Polygenic Risk Scores (PRS) are usually evaluated against the incidence of diseases they were derived for. Individuals may be more interested in how these PRS influence their probable cause of death. Using UK Biobank data, we examined the top 10 causes of death among individuals in the highest quintile of disease-specific PRS for Alzheimers disease, bowel cancer, cardiovascular disease, coronary artery disease, ischaemic stroke, breast cancer, epithelial ovarian cancer, and prostate cancer. Analyses were stratified by sex, age at death, and smoking status (never, past, current). We also assessed varying PRS percentile thresholds to identify when the target disease became the leading cause of death, and evaluated the impact of each disease-specific PRS on all-cause mortality using Cox proportional hazards models. For most disease-specific PRS, individuals in the high-risk group were more likely to die from other common diseases. The leading causes of death varied according to demographic and behavioural subgroup: breast cancer in women, ischaemic heart disease in men, dementia in the oldest age groups, and lung cancer among smokers. For instance, while prostate cancer was the leading cause of death among older never-smoking men in the highest quintile of the prostate cancer PRS; in other age and smoking status categories, ischaemic heart disease or lung cancer were more common. While a high PRS is predictive of disease diagnosis, most individuals die from other common conditions, depending on their demographic and behavioural subgroups. These findings highlight the importance of contextualising PRS results in clinical settings and risk communication. 2. Key messagesO_LIWhat is already known on this topic: C_LI Disease-specific PRS have been investigated for their ability to predict incidence, not death, from the specific target disease. O_LIWhat this study adds: C_LI We evaluated PRS for the most common diseases against death from the target disease, as well as other common causes of death. O_LIHow this study might affect research, practice or policy: C_LI Providing the probabilities of death from each target disease, and from other diseases, to the probability of PRS-specific incidence may help contextualise communication of risks associated with high disease-specific PRS.
Qi, X.; Qi, H.; li, N.; Wang, T.; Wang, W.; Song, X.; Mi, B.; Zhang, D.
Show abstract
ABSTRACT Background and aims: Mental and behavioral disorders due to use of tobacco (MBDT) present a critical challenge to global health, yet modifiable lifestyle factors for reducing its risk remain poorly understood. Given that dietary fibre can affect mental health through gut-brain communication, we sought to explore how fibre intake relates to MBDT risks in smokers. Methods: We specifically evaluated the link between dietary fibre intake and MBDT within a smoking population. Utilizing the UK Biobank (UKB) database, we performed cross-sectional (N=19,943) and prospective cohort (N=19,885) evaluations applying logistic and Cox proportional hazards models, respectively. To determine potential causality, two-sample Mendelian randomization (MR) was applied, relying on GWAS summary data derived from the IEU Open GWAS Project and FinnGen repositories. Results: Cross-sectional findings indicated that individuals in the top quartile (Q4) of fibre intake exhibited decreased MBDT risks relative to the bottom quartile (Q1) (OR: 0.32, 95% CI: 0.13-0.79). Over a median observation time of 12.84 years, the prospective evaluation demonstrated a notable inverse correlation (Q4 HR: 0.46, 95% CI: 0.40-0.54). Non-linear modeling via restricted cubic splines uncovered an L-shaped dose-response curve. Furthermore, MR results confirmed a genetically predicted protective causality (IVW OR: 0.68, 95% CI: 0.49-0.95), which remained consistent across sensitivity validations. Conclusions: Among smokers, higher dietary fibre intake is robustly associated with a reduced risk of mental and behavioral disorders due to the use of tobacco, offering a modifiable dietary target for public health interventions.
Ahlqvist, V. H.; Sjoqvist, H.; Sjolander, A.; Berglind, D.; Lambert, P. C.; Lee, B. K.; Madley-Dowd, P.
Show abstract
ObjectiveFindings from family-based analyses, such as sibling comparisons, are often reported using only odds ratios or hazard ratios. We demonstrate how this can be improved upon by applying the marginalized between-within framework. Study Design and SettingWe provide an overview of sibling comparison methods and the marginalized between-within framework, which enables estimation of absolute risks and clinically relevant metrics while accounting for shared familial confounding. We illustrate the approach using Swedish registry data to examine the association between maternal smoking and infant mortality, estimating absolute risk differences, average treatment effects, attributable fractions, and numbers needed to harm (or treat). ResultsThe marginalized between-within model decomposes effects into within-and between-family components while applying a global baseline across all families. Although it typically yields similar relative estimates to conditional logistic or stratified Cox regression, the models specification of a baseline enables the estimation of absolute measures. In the applied example, absolute measures provided more interpretable and policy-relevant insights than relative estimates alone. Code for implementation in Stata and R is provided. ConclusionThe marginalized between-within framework may strengthen the interpretability of family-based analysis by enabling absolute and policy-relevant estimates for both binary and time-to-event outcomes, moving beyond the limitations of solely relying on relative effect measures. What is new?O_ST_ABSKey FindingsC_ST_ABSO_LIFindings from sibling analyses are typically presented using only relative measures, such as odds ratios or hazard ratios, limiting interpretability. C_LIO_LIThis study illustrates how the marginalized between-within framework can be used to derive clinically relevant absolute effect measures while adjusting for shared familial confounding. C_LI What this adds to what was known?O_LIUnlike conventional methods, this approach enables estimation of absolute risks, average treatment effects, attributable fractions, and numbers needed to treat or harm--using standard software--while accounting for unmeasured familial con-founding. C_LI What is the implication and what should change now?O_LIResearchers conducting sibling comparisons should consider adopting the marginalized between-within framework to report both relative and absolute effect measures. C_LIO_LIThis shift could enhance the clinical and public health relevance of family-based designs by improving interpretability and communication of findings. C_LI
Taylor, K.; Howe, L. D.; Lacey, R. E.; Carslake, D.; Anderson, E. L.; Mukadam, N.
Show abstract
IntroductionStudies investigating the association between adverse experiences across the life-course and dementia consider a narrow range of experiences and use sum scores, assuming each experience has the same impact on dementia risk. We considered the timing, type and cumulation of adverse experiences. MethodsThe English Longitudinal Study of Ageing measured adverse experiences in a retrospective interview. Cox proportional hazard models were used to investigate associations between dementia and sum adversity scores, individual experiences, and broad categories adapted from existing frameworks. ResultsNumber of adult, but not total or childhood, adverse experiences was associated with dementia. Child abuse and adult economic hardship were associated with a 74% and 32% higher hazard of dementia respectively. DiscussionAdulthood adverse experiences associate with dementia in a cumulative risk manner. In childhood, only abuse was associated with dementia. Use of sum scores to summarise adverse experiences throughout the life-course may oversimplify associations with dementia.
Riedmann, U.; Levitt, M.; Pilz, S.; Ioannidis, J.
Show abstract
BackgroundPost-pandemic mortality rates can explore the residual COVID-19 burden and changes in other causes of death. Considering weighted multiple causes of death from death certificates (underlying and others) may help compare post-versus pre-pandemic mortality patterns, while potentially reducing the impact of cause misattribution. Estimates of post-pandemic impact are critical also for proper continuing public health policies (e.g. vaccinations). MethodsWe retrospectively analyse national all-cause mortality rate ratios between 2024 and pre-pandemic years (2017-2019) for sex-stratified 10-year age groups in Austria. In weighted analyses, the underlying death cause was weighted 50% and other causes shared the remaining 50%. Sensitivity analyses explored different weightings. Cause-specific weightings were also compared between 2024 and 2019. ResultsDespite 1,212 reported COVID-19 deaths in 2024, all-cause mortality rates were equal or lower in 2024 compared to 2019 in all strata at risk from COVID-19 (i.e., aged 60 years and over). All-cause mortality rates in 2024 were higher than in 2019 in adolescent and young adult strata. The ratio of weighted over unweighted COVID-19 death rates was 0.51-0.58 for age strata 60 years and older and even lower in sensitivity analyses, indicating that COVID-19 deaths were possibly overestimated. ConclusionsPost-pandemic COVID-19 deaths had no visible impact on mortality patterns in Austria and were possibly overcounted. Increased post-pandemic mortality patterns in the young are particularly worrisome.
Jee, Y.; Spiller, W.; Sanderson, E.; Tilling, K.; Palmer, T.; Ha, E.; Kim, Y. J.
Show abstract
This study evaluates the potential role of multiple correlated risk factors upon coronary heart disease (CHD) and ischemic stroke, and the extent to which using GWAS summary data including prevalent cases of stroke, as opposed to incident cases, can influence Mendelian randomization (MR) analyses. Initially, thirteen candidate risk factors were identified through a literature review, including age of menarche, adiposity, blood pressure, lipid fractions, physical activity, type-II diabetes, smoking, sleep duration, alcohol consumption, and kidney function. Using publicly available summary data from genome-wide association studies (GWAS), the total effect of each exposure on CHD, ischemic, and cardioembolic stroke was estimated using univariable summary MR. Multivariable MR (MVMR) analyses were then used to estimate the conditional effects of low-density lipoprotein (LDL), high-density lipoprotein (HDL), triglycerides and systolic blood pressure (SBP) on each outcome. To select the MVMR model a novel forward selection algorithm was applied to include the greatest number of exposures while maintaining sufficient conditional instrument strength for estimation. To examine potential bias from using GWAS summary data derived from prevalent cases of ischemic stroke a GWAS of incident ischemic stroke was conducted using data from the UK Biobank. In univariable MR analyses negative effects of blood pressure were observed across all outcomes, while the effects of remaining exposures differed markedly. HDL was also estimated to have a protective effect on all outcomes except cardioembolic stroke. Univariable and MVMR estimates were directionally consistent, though MVMR estimates were attenuated. Finally, repeating analyses using incident stroke cases yielded results in agreement with prevalent stroke data, suggesting the use of prevalent outcome data did not bias our initial analysis.
Herranen, P.; Koivunen, K.; Palviainen, T.; Finn gen, ; Kujala, U. M.; Ripatti, S.; Kaprio, J.; Sillanpää, E.
Show abstract
PurposeTo use a genome-wide polygenic risk score for hand grip strength (PRS HGS) to investigate whether the muscle strength genotype predicts the most common age related noncommunicable diseases, survival from acute adverse health events, and all cause mortality. MethodsThis study consisted of 342 443 Finnish biobank participants from FinnGen Data Freeze 10 (53% women) aged 40 to 108 with combined genotype and health registry data. Associations were explored with a linear or Cox proportional hazards regression models. ResultsA higher PRS HGS predicted a lower body mass index (BMI) ({beta} = -0.112 kg/m2, standard error (SE) = 0.017, P = 1.69 x 10-11) in women but not in men ({beta} = 0.004 kg/m2, P = 0.768, sex by PRS HGS interaction: P = 2.12 x 10-07). In all participants, a higher PRS HGS was associated with a lower risk for obesity diagnosis (hazard ratio 0.94, 95% confidence interval 0.93 to 0.95), type 2 diabetes (0.95, 0.94 to 0.96), ischemic heart diseases (0.97, 0.96 to 0.97), hypertension (0.97, 0.96 to 0.97), stroke (0.97, 0.96 to 0.98), asthma (0.94, 0.93 to 0.95), chronic obstructive pulmonary disease (0.94, 0.92 to 0.95), polyarthrosis (0.90, 0.88 to 0.92), knee arthrosis (0.98, 0.97 to 0.99), rheumatoid arthritis (0.95, 0.94 to 0.97), osteoporosis (0.95, 0.93 to 0.97), falls (0.98, 0.98 to 0.99), depression (0.95, 0.94 to 0.96), and vascular dementia (0.93, 0.89 to 0.96). In women only, a higher PRS predicted a lower hazard for any dementia (0.94, 0.92 to 0.96) and Alzheimers disease (0.96, 0.93 to 0.98). Participants with a higher PRS HGS had a decreased risk of cardiovascular (0.96, 0.95 to 0.98) and all cause mortality (0.97, 0.96 to 0.98). However, the predictive value of the PRS HGS for mortality was not pronounced after adverse acute health events compared to the non-diseased period. ConclusionsThe genotype that supports higher muscle strength protects against many future health adversities. Further research is needed to investigate whether or how a favourable lifestyle modifies this intrinsic capacity to resist diseases, and if the impacts of lifestyle behaviour on health differ due to polygenic risk.
Pezzullo, A. M.; Axfors, C.; Contopoulos-Ioannidis, D. G.; Apostolatos, A.; Ioannidis, J. P. A.
Show abstract
The infection fatality rate (IFR) of COVID-19 among non-elderly people in the absence of vaccination or prior infection is important to estimate accurately, since 94% of the global population is younger than 70 years and 86% is younger than 60 years. In systematic searches in SeroTracker and PubMed (protocol: https://osf.io/xvupr), we identified 40 eligible national seroprevalence studies covering 38 countries with pre-vaccination seroprevalence data. For 29 countries (24 high-income, 5 others), publicly available age-stratified COVID-19 death data and age-stratified seroprevalence information were available and were included in the primary analysis. The IFRs had a median of 0.035% (interquartile range (IQR) 0.013 - 0.056%) for the 0-59 years old population, and 0.095% (IQR 0.036 - 0.125%,) for the 0-69 years old. The median IFR was 0.0003% at 0-19 years, 0.003% at 20-29 years, 0.011% at 30-39 years, 0.035% at 40-49 years, 0.129% at 50-59 years, and 0.501% at 60-69 years. Including data from another 9 countries with imputed age distribution of COVID-19 deaths yielded median IFR of 0.025-0.032% for 0-59 years and 0.063-0.082% for 0-69 years. Meta-regression analyses also suggested global IFR of 0.03% and 0.07%, respectively in these age groups. The current analysis suggests a much lower pre-vaccination IFR in non-elderly populations than previously suggested. Large differences did exist between countries and may reflect differences in comorbidities and other factors. These estimates provide a baseline from which to fathom further IFR declines with the widespread use of vaccination, prior infections, and evolution of new variants. Highlights*Across 31 systematically identified national seroprevalence studies in the pre-vaccination era, the median infection fatality rate of COVID-19 was estimated to be 0.035% for people aged 0-59 years people and 0.095% for those aged 0-69 years. *The median IFR was 0.0003% at 0-19 years, 0.003% at 20-29 years, 0.011% at 30-39 years, 0.035% at 40-49 years, 0.129% at 50-59 years, and 0.501% at 60-69 years. *At a global level, pre-vaccination IFR may have been as low as 0.03% and 0.07% for 0-59 and 0-69 year old people, respectively. *These IFR estimates in non-elderly populations are lower than previous calculations had suggested.
Arning, N.; Fryer, H. R.; Wilson, D. J.
Show abstract
Big data approaches to discovering non-genetic risk factors have lagged behind genome-wide association studies that routinely uncover novel genetic risk factors for diverse diseases. Instead, epidemiology typically focuses on candidate risk factors. Since modern biobanks contain thousands of potential risk factors, candidate approaches may introduce bias, inadequately control for multiple testing, and overlook important signals. Doublethink, a novel model-averaged hypothesis testing approach, offers a solution that simultaneously controls the Bayesian false discovery rate (FDR) and frequentist familywise error rate (FWER) while accounting for uncertainty in variable selection. Here we investigate direct risk factors for COVID-19 hospitalization from among 1,912 variables in 201,917 UK Biobank participants by implementing a Doublethink-based exposome-wide association study using Markov Chain Monte Carlo. Focusing on the 2020 outbreak, we find nine individual variables and six groups of variables exposome-wide significant at 9% FDR and 0.05% FWER. We identify significant direct effects among relatively overlooked risk factors including psychiatric disorders, dementia and prior infection, which we evaluate in relation to studies of other populations. We detect significant direct effects among some commonly reported risk factors like age, sex and obesity, but not others like diabetes, cardiovascular disease, hypertension, which may be mediated instead through variables representing general comorbidity. Doublethink produces interchangeable posterior odds and p-values for individual variables and arbitrary groups, facilitating flexible and powerful post-hoc hypothesis testing. We discuss the potential for impact and limitations of joint Bayesian-frequentist hypothesis testing, including the benefits of an agnostic exposome-wide approach to discovery. SignificanceUnderstanding what causes disease is key to improving its treatment and prevention. Large health studies like UK Biobank measure thousands of possible causes of disease. Traditionally, scientists have studied possible causes (like smoking or exercise) one-at-a-time, in depth. For greater perspective, we could study them altogether to test which have any effect. We recently introduced Doublethink, which combines the advantages of two major statistical approaches to testing. Here we use Doublethink to test 1,912 possible causes of COVID-19 hospitalization in UK Biobank. We found strong evidence for relatively overlooked causes: psychiatric conditions, dementia and previous infections. Findings from other health studies support these causes, highlighting the need to re-evaluate them and showing how our approach can reveal valuable insights.
Jones, P.; Bhatta, L.; Howe, L.; Vinueza-Veloz, M. F.; Davey Smith, G.; Naess, O. E.; Brumpton, B. M.
Show abstract
BackgroundObservational studies have consistently found educational inequalities in cardiovascular disease risk. Mendelian randomisation analyses have suggested a direct causal effect of education, however estimates may be biased by demography or dynastic effects. This study aimed to estimate the effects of educational attainment on cardiovascular disease risk and serum lipid concentrations before and after accounting for family structure. MethodsThis study included 28 448 siblings from the Trondelag Health Study (HUNT), 27 040 siblings from UK Biobank, and >120 000 individuals from an international within-sibship genome-wide association study, predominantly of European ancestry. The exposure was educational attainment. The outcomes were cardiovascular disease risk and serum concentrations of low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, and triglycerides. Standard and within-sibship Mendelian randomisation were used. ResultsIn the summary data analysis, there was a 6% lower risk of cardiovascular disease (odds ratio 0.94, 95% confidence interval 0.92 to 0.96) for each additional standard deviation of educational attainment. This was consistent having accounted for family structure (odds ratio 0.96, 95% confidence interval 0.91 to 1.01). Educational attainment was also beneficially associated with each serum lipid concentration both before and after accounting for family structure. Results were broadly similar in the individual-level analysis. ConclusionsThere is a protective effect of educational attainment on cardiovascular disease risk and a beneficial effect on serum lipid concentrations not due to familial factors shared by siblings, suggesting that increasing education may be beneficial for cardiovascular health. KEY MESSAGESO_ST_ABSQuestionC_ST_ABSIs the direct causal effect of education on cardiovascular disease (CVD) risk and CVD risk factors indicated by conventional Mendelian randomisation (MR) biased by demography or dynastic effects? FindingsConsistent with conventional MR analyses, within-sibship MR indicated that higher educational attainment is protective CVD risk and beneficial for serum concentrations of low-density lipoprotein cholesterol, high-density lipoprotein cholesterol, and triglycerides. ImportancePreviously reported beneficial effects of educational attainment on CVD risk and serum lipid concentrations are likely causal and not due to bias from demographic or dynastic effects.
Chudasama, Y. V.; Khunti, K.; Gillies, C. L.; Dhalwani, N. N.; Davies, M. J.; Yates, T.; Zaccardi, F.
Show abstract
Background and objectiveThere has been an increasing interest in using life expectancy metrics, such as years of life lost (YLL), to explore epidemiological associations. YLL is easier to understand for both healthcare professionals and the lay people and has become a common measure for evaluating public health priorities. As the literature presents a range of approaches to estimate it, this review aims to: (1) summarise the key methods; (2) show how to implement them using current software; (3) apply them in a real-world example. MethodsWe investigated simpler nonparametric as well as parametric, model-based methods to estimate of YLL, including: (1) Years of potential life lost (YPLL); (2) Global Burden of Disease (GBD) approach; (3) Chiangs life tables; (4) Epi-demographic approach; and (5) Flexible Royston-Parmar parametric survival model. We used data from the UK Biobank with baseline measures collected in 2006-2010 and linkage to mortality records. We selected 36 chronic conditions: participants with two or more conditions were categorised as having multimorbidity. ResultsFor the YPLL and GBD method, the analytical procedures allow only to quantify the average YLL within each group (with and without multimorbidity) and, from them, their difference. Conversely, for the Chiangs life tables, the epi-demographic approach, and the Royston-Parmar survival model, both the remaining life expectancy within each group and the YLL could be estimated. In 499,992 UK Biobank participants (white ethnicity, 94%; women, 55%) with a median (IQR) age of 58 (50-63) years, 98,605 (20%) had multimorbidity and 11,871 deaths occurred during the follow-up. The YLLs comparing subjects with vs without multimorbidity varied significantly according to the technique and the modelling approach used: from a longer life expectancy in subjects with multimorbidity using the YPLL and the GBD method to a shorter one using the other three methods (i.e., at 65 years, the YLL were 1.8, 1.3, and 4.6 years using Chiangs, epi-demographic, and Royston-Parmar approach, respectively). ConclusionsWhen comparing the burden of a disease on life expectancy across studies caution is needed as methods may estimate different quantities. While deciding among different methods to estimate YLL, researchers should consider such differences in relation to the purpose of the research and the type of available data. O_TEXTBOXSUMMARY BOXO_LIThe concept of years of life lost (YLL) is easier to understand compared to traditional estimates from survival analysis, such as hazard ratios, but very few studies report it. C_LIO_LIA range of different methods of estimating YLL are reviewed: from basic methods - such as life tables - to most recent and advanced methods using statistical modelling. C_LIO_LIUsing the example of multimorbidity, the estimated numbers of YLL differs between methods, as each method focused on estimating different quantities. C_LIO_LIThis review will help promote a better understanding and use of life expectancy and YLL metrics in a wide range of studies in health care research. C_LI C_TEXTBOX
Millard, L. A. C.; Fernandez-Sanles, A.; Carter, A. R.; Hughes, R.; Tilling, K.; Morris, T. P.; Smith, D.; Griffith, G. J.; Clayton, G. L.; Kawabata, E.; Davey Smith, G.; Lawlor, D. A.; Borges, M. C.
Show abstract
BackgroundNon-random selection into analytic subsamples could introduce selection bias in observational studies of SARS-CoV-2 infection and COVID-19 severity (e.g. including only those have had a COVID-19 PCR test). We explored the potential presence and impact of selection in such studies using data from self-report questionnaires and national registries. MethodsUsing pre-pandemic data from the Avon Longitudinal Study of Parents and Children (ALSPAC) (mean age=27.6 (standard deviation [SD]=0.5); 49% female) and UK Biobank (UKB) (mean age=56 (SD=8.1); 55% female) with data on SARS-CoV-2 infection and death-with-COVID-19 (UKB only), we investigated predictors of selection into COVID-19 analytic subsamples. We then conducted empirical analyses and simulations to explore the potential presence, direction, and magnitude of bias due to selection when estimating the association of body mass index (BMI) with SARS-CoV-2 infection and death-with-COVID-19. ResultsIn both ALSPAC and UKB a broad range of characteristics related to selection, sometimes in opposite directions. For example, more educated participants were more likely to have data on SARS-CoV-2 infection in ALSPAC, but less likely in UKB. We found bias in many simulated scenarios. For example, in one scenario based on UKB, we observed an expected odds ratio of 2.56 compared to a simulated true odds ratio of 3, per standard deviation higher BMI. ConclusionAnalyses using COVID-19 self-reported or national registry data may be biased due to selection. The magnitude and direction of this bias depends on the outcome definition, the true effect of the risk factor, and the assumed selection mechanism. Key messagesO_LIObservational studies assessing the association of risk factors with SARS-CoV-2 infection and COVID-19 severity may be biased due to non-random selection into the analytic sample. C_LIO_LIResearchers should carefully consider the extent that their results may be biased due to selection, and conduct sensitivity analyses and simulations to explore the robustness of their results. We provide code for these analyses that is applicable beyond COVID-19 research. C_LI
Yu, X.; Lophatananon, A.; Holmes, V.; Muir, K. R.; Guo, H.
Show abstract
INTRODUCTIONComprehensively studying modifiable risk factors altogether to explore how they contribute to dementia mechanism is imperative for effective interventions. METHODSThis study utilized natural language processing (NLP) models to pre-select candidate risk factors of dementia from 5,505 variables in the UK Biobank. We then took a holistic machine learning approach, fast causal inference in combination with mixed graphical models, to explore the complex causal mechanisms underlying dementia from 120 imputed variables. RESULTSThe identified causal network highlighted eight risk factors which may directly or indirectly contribute to dementia. In particular, mental disorders due to brain damage and dysfunction and to physical disease were identified as direct causes as well as mediators on the causal pathways to dementia. Evidence for a direct causal impact of phenotypic age on dementia was less pronounced. DISCUSSIONOur study offered valuable insights into the mechanisms of dementia. Beyond direct connections to nerve or brain disorders, the potential direct link with biological age highlights its possible value in dementia management. Moreover, the use of NLP models for variable pre-selection introduced an innovative application to medical research. Our study added weight to the accruing evidence that machine learning has a promising future for exploring complex disease mechanisms from high-dimensional data.
Taylor, K.; Howe, L. D.; Lacey, R.; Anderson, E. L.; Mukadam, N.
Show abstract
Background Literature investigating mediation of the association between child abuse and dementia has largely considered composite adverse childhood experience scores rather than individual adverse experiences, despite evidence that different experiences have different impacts on dementia risk. Additionally, prior studies consider mediators in isolation, despite known associations between mediators which may impact indirect pathways from child abuse to dementia. Objectives To investigate whether potentially modifiable health and lifestyle factors mediate the association between child abuse and dementia. Methods We used data from the English Longitudinal Study of Ageing to investigate associations between child abuse and dementia (N:5,448). Indirect pathways through four mediator categories (education, health behaviours, mental health and cardiovascular health) were examined. We used regression modelling to estimate associations between child abuse, mediators and dementia, and causal mediation analysis using the g-formula to estimate the joint indirect effect through the mediators. Results Individuals who experienced child abuse had, on average, an 80% higher hazard of dementia, compared to those who did not (RTE HR:1.80, 95% CI:1.21-2.39). Mental health mediators showed strong associations with both child abuse and dementia. Evidence for other mediators was weaker. Education, health behaviours, mental health and cardiovascular health mediated approximately 18% of the association. Sensitivity analysis revealed that almost all this mediation occurred through mental health. Conclusions Child abuse was associated with higher risk of dementia. Joint mediation analysis suggested that education, health behaviours, cardiovascular health, and mental health accounted for a relatively small proportion of the observed association, with most mediation occurring through mental health. Future research must focus on other potential pathways from child abuse to dementia, including biological and social mechanisms.
Webster, A. J.
Show abstract
Epidemiologists are careful to describe their findings as "associations", and to avoid any causal language or claims. Arguably, this attempt to avoid reference to causal processes has become counterproductive. Explicitly stated or not, assumptions about causal processes are inherent in the formulation and interpretation of any statistical study. This article offers a bridge between established, extensively developed proportional hazard methods that are used to study longitudinal observational cohort data, and results for causal inference. In particular, it considers the burden of disease that would not have occurred, but for an exposure such as smoking. It shows how this "probability of necessity", relates to population attributable fractions, and how these quantities along with their confidence intervals, can be estimated using conventional proportional hazard estimates. The example may often apply to cohort studies that consider disease-risk in the absence of prior disease. More generally, equivalent estimates can often be constructed when there is sufficient understanding to postulate a model for the causal relationship between exposures, confounders, and disease-risk, as summarised in a directed acyclic graph (DAG).
Katsoulis, M.; Narayanan, M.; Dodgeon, B.; Ploubidis, G.; Silverwood, R.
Show abstract
BackgroundMissing data may induce bias when analysing longitudinal population surveys. We aimed to tackle this problem in the 1970 British Cohort Study (BCS70) MethodsWe utilised a data-driven approach to address missing data issues in BCS70. Our method consisted of a 3-step process to identify important predictors of non-response from a pool of [~]20,000 variables from 9 sweeps in 18037 individuals. We used parametric regression models to identify a moderate set of variables (predictors of non-response) that can be used as auxiliary variables in principled methods of missing data handling to restore baseline sample representativeness. ResultsIndividuals from disadvantaged socio-economic backgrounds, increased number of older siblings, non-response at previous sweeps and ethnic minority background were consistently associated with non-response in BCS70 at both early (ages 5-16) and later sweeps (ages 26-46). Country of birth, parents not being married and higher fathers age at completion of education were additional consistent predictors of non-response only at early sweeps. Moreover, being male, greater number of household moves, low cognitive ability, and non-participation in the UK 1997 elections were additional consistent predictors of non-response only at later sweeps. Using this information, we were able to restore sample representativeness, as we could replicate the original sample distribution of fathers social class and cognitive ability and reduce the bias due to missing data in the relationship between fathers socioeconomic status and mortality. ConclusionsWe provide a set of variables that researchers can utilise as auxiliary variables to address missing data issues in BCS70 and restore sample representativeness. Key MessagesO_LIWe aimed to address the problem of missing data in the 1970 British Cohort Study (BCS70) caused by non-response at different sweeps C_LIO_LIWe identified a set of predictors of non-response that can successfully restore baseline sample representativeness across sweeps C_LIO_LIThe information from this study can be used from researchers in the future to utilise appropriate auxiliary variables to tackle problems due to missing data in BCS70 C_LI
Batty, G. D.; Gale, C.; Kivimaki, M.; Deary, I.; Bell, S.
Show abstract
BackgroundThe UK Biobank cohort study has become a much-utilised and influential scientific resource. With a primary goal of understanding disease aetiology, the low response to the original survey of 5.5% has, however, led to debate as to the generalisability of these findings. We therefore compared risk factor-disease estimations in UK Biobank with those from 18 nationally representative studies with conventional response rates. MethodsWe used individual-level baseline data from UK Biobank (N=502,655) and a pooling of data from the Health Surveys for England (HSE) and the Scottish Health Surveys (SHS), comprising 18 studies and 89,895 individuals (mean response rate 68%). Both study populations were aged 40-69 years at study induction and linked to national cause-specific mortality registries. FindingsDespite a typically more favourable risk factor profile and lower mortality rates in UK Biobank participants relative to the HSE-SHS consortium, risk factors-endpoints associations were directionally consistent between studies, albeit with some heterogeneity in magnitude. For instance, for cardiovascular disease mortality, the age- and sex-adjusted hazard ratio (95% confidence interval) for ever having smoked cigarettes (versus never) was 2.04 (1.87, 2.24) in UK Biobank and 1.99 (1.78, 2.23) in HSE-SHS, yielding a ratio of hazard ratios close to unity (1.02, 0.88, 1.19; p-value 0.76). For hypertension (versus none), corresponding results were again in same direction but with a lower effect size in UK Biobank (1.89; 1.69, 2.11) than in HSE-SHS (2.56; 2.20, 2.98), producing a ratio of hazard ratios below unity (0.74; 0.62, 0.89; p-value 0.001). A similar pattern of observations were made for risk factors (smoking, obesity, educational attainment, and physical stature) in relation to different cancer presentations and suicide whereby the ratios of hazard ratios ranged from 0.57 (0.40, 0.81) and 1.07 (0.42, 2.74). InterpretationDespite a low response rate, aetiological findings from UK Biobank appear to be generalisable to England and Scotland.
de Lange, M. A.; Davies, N. M.; Millard, L. A. C.; Tilling, K.
Show abstract
BackgroundObservational research shows that a childs relative age within their school year ( relative age) is associated with educational attainment and mental health. However, previous studies have only examined a small number of outcomes and evidence of the persistence of effects into adulthood is mixed. We conducted a hypothesis-free investigation of the effects of relative age. MethodWe used a regression discontinuity design and an instrumental variable (IV)-pheWAS in the UK Biobank (participants aged 40-69 years at baseline), using the PHESANT software package. We created two IVs for relative age: being born in September vs. August (n=64 075) and week of birth (n=383 309). Outcomes passing the Bonferroni-corrected P value threshold for either instrument were plotted to identify those displaying a discontinuity at the school year transition. ResultsWe found 21 traits associated with at least one of the instruments (P value below the Bonferroni threshold). Of these, 13 showed a discontinuity at the school year transition. These included previously identified effects including those with a younger relative age being less likely to have educational qualifications and more likely to have started smoking at an earlier age. We also identified a novel potential effect of a younger relative age in school year causing a better lung function as adults. ConclusionEducational policy should address educational inequality due to relative age. Further research should seek to replicate our identified effect on lung function in different populations, and investigate the mechanisms through which this effect may act. Key MessagesO_LIChildrens relative age within their school year has been associated with mental health in childhood and educational attainment. C_LIO_LIOur results supported previously identified effects, with those who were younger in their school year being less likely to have educational qualifications and more likely to report starting smoking at an earlier age. C_LIO_LIWe also found a potential beneficial effect of a younger relative age in school year on lung function in adulthood. C_LI
Bonnet, F.; Klusener, S.; Mesle, F.; Muhlichen, M.; Grigoriev, P.
Show abstract
BackgroundBoth enhancing life expectancy as well as diminishing inequalities in lifespan among social groups represent significant goals for public policy. However, there is a lack of methodological tools to simultaneously monitor progress in both dimensions. Additionally, there is a consensus that absolute and relative inequalities in lifespan must be scrutinized together. MethodsWe introduce a novel graphical representation that combines national mortality rates with social inequalities, considering both absolute and relative measures. We use French and German data stratified by place of residence to illustrate this representation. ResultsFor all-age mortality we detect for France a rather continuous pace of decline in both mortality levels and variation. In Germany, substantial progress was made in the 1990s, which was mostly driven by convergence between eastern and western Germany, followed by a period with less progress. Age-specific analyses reveal for Germany some worrying regional divergence trends at ages 35-74 in recent years. This is particularly pronounced among women. ConclusionOur novel visual approach allows evaluating easily the dynamics of societal progress in terms of longevity, and facilitates meaningful comparisons between populations, even when their current mortality rates differ. The methods we employ can be reproduced easily in any country with longitudinal mortality data stratified by relevant socio-economic information or regions. It is both useful for scientific analyses as well as policy advice. Key messagesO_ST_ABSWhat is already known on this topicC_ST_ABSImproving life expectancy as well as reducing social inequalities in longevity are major public policy objectives. However, there is a lack of proper methodological tools to evaluate progress on these objectives. What this study addsThis study proposes an innovative graphical representation that combines national mortality and social inequalities in both absolute and relative terms in order to assess the dynamics of societal progress in longevity and make relevant comparisons between populations whose mortality rates are not at the same level nowadays. How this study might affect research, practice or policyMethods are freely and easily reproducible for all countries with longitudinal mortality data stratified by socio-economic information or geographic regions.
Clayton, G. L.; Goncalves Soares, A.; Goulding, N.; Borges, M. C. G.; Holmes, M.; Davey Smith, G.; Tilling, K.; Lawlor, D. A.; Carter, A. R.
Show abstract
ObjectiveTo use the example of the effect of body mass index (BMI) on COVID-19 susceptibility and severity to illustrate methods to explore potential selection and misclassification bias in Mendelian randomisation (MR) of COVID-19 determinants. DesignTwo-sample MR analysis. SettingSummary statistics from the Genetic Investigation of ANthropometric Traits (GIANT) and COVID-19 Host Genetics Initiative (HGI) consortia. Participants681,275 participants in GIANT and more than 2.5 million people from the COVID-19 HGI consortia. ExposureGenetically instrumented BMI. Main outcome measuresSeven case/control definitions for SARS-CoV-2 infection and COVID-19 severity: very severe respiratory confirmed COVID-19 vs not hospitalised COVID-19 (A1) and vs population (those who were never tested, tested negative or had unknown testing status (A2)); hospitalised COVID-19 vs not hospitalised COVID-19 (B1) and vs population (B2); COVID-19 vs lab/self-reported negative (C1) and vs population (C2); and predicted COVID-19 from self-reported symptoms vs predicted or self-reported non-COVID-19 (D1). ResultsWith the exception of A1 comparison, genetically higher BMI was associated with higher odds of COVID-19 in all comparison groups, with odds ratios (OR) ranging from 1.11 (95%CI: 0.94, 1.32) for D1 to 1.57 (95%CI: 1.57 (1.39, 1.78) for A2. As a method to assess selection bias, we found no strong evidence of an effect of COVID-19 on BMI in a no-relevance analysis, in which COVID-19 was considered the exposure, although measured after BMI. We found evidence of genetic correlation between COVID-19 outcomes and potential predictors of selection determined a priori (smoking, education, and income), which could either indicate selection bias or a causal pathway to infection. Results from multivariable MR adjusting for these predictors of selection yielded similar results to the main analysis, suggesting the latter. ConclusionsWe have proposed a set of analyses for exploring potential selection and misclassification bias in MR studies of risk factors for SARS-CoV-2 infection and COVID-19 and demonstrated this with an illustrative example. Although selection by socioeconomic position and arelated traits is present, MR results are not substantially affected by selection/misclassification bias in our example. We recommend the methods we demonstrate, and provide detailed analytic code for their use, are used in MR studies assessing risk factors for COVID-19, and other MR studies where such biases are likely in the available data. SummaryO_ST_ABSWhat is already known on this topicC_ST_ABS- Mendelian randomisation (MR) studies have been conducted to investigate the potential causal relationship between body mass index (BMI) and COVID-19 susceptibility and severity. - There are several sources of selection (e.g. when only subgroups with specific characteristics are tested or respond to study questionnaires) and misclassification (e.g. those not tested are assumed not to have COVID-19) that could bias MR studies of risk factors for COVID-19. - Previous MR studies have not explored how selection and misclassification bias in the underlying genome-wide association studies could bias MR results. What this study adds- Using the most recent release of the COVID-19 Host Genetics Initiative data (with data up to June 2021), we demonstrate a potential causal effect of BMI on susceptibility to detected SARS-CoV-2 infection and on severe COVID-19 disease, and that these results are unlikely to be substantially biased due to selection and misclassification. - This conclusion is based on no evidence of an effect of COVID-19 on BMI (a no-relevance control study, as BMI was measured before the COVID-19 pandemic) and finding genetic correlation between predictors of selection (e.g. socioeconomic position) and COVID-19 for which multivariable MR supported a role in causing susceptibility to infection. - We recommend studies use the set of analyses demonstrated here in future MR studies of COVID-19 risk factors, or other examples where selection bias is likely.